Papers with computational models
Copied to clipboard
| Challenge: | Prior work has explored the ability of computational models to predict word semantic fit with a given predicate. |
| Approach: | They compare Transformers Language Models to SDM to assess their performance . they found that TLMs do not capture important aspects of event knowledge . people can discriminate between typical and atypical events, they say . |
| Outcome: | The proposed models can achieve comparable performance to SDM, but they lack important aspects of event knowledge. |
Copied to clipboard
| Challenge: | a tutorial focuses on computational models for conversational structure, summarization and sentiment detection, and group dynamics. |
| Approach: | a tutorial will provide examples of specific NLP tasks for conversational structure, summarization and sentiment detection, and group dynamics. |
| Outcome: | The tutorial focuses on the three areas of conversational structure, summarization and sentiment detection, and group dynamics. |
Copied to clipboard
| Challenge: | Argumentation is a rhetorical device that asserts propositions implicitly, but few studies have examined the issue. |
| Approach: | They propose a computational method for extracting propositions that are implicitly asserted in questions, reported speech, and imperatives in argumentation. |
| Outcome: | The proposed models are based on a corpus of 2016 debates and online commentary. |
Copied to clipboard
| Challenge: | Multilingual communities exhibit code-mixing, mixing of two or more languages in a single conversation . social media and other informal interactive platforms are allowing code-switching in user-generated text . |
| Approach: | a tutorial aims to provide a foundation for researchers to study code-mixing in multilingual communities. |
| Outcome: | a tutorial aims to provide new researchers with a foundation in linguistics and computational aspects of code-mixing. |
Copied to clipboard
| Challenge: | Chinese characters encode world knowledge through thousands of years evolution . |
| Approach: | They propose an embedding approach to encode Chinese orthography knowledge using eigencharacter space. |
| Outcome: | The proposed representations encode lexical knowledge embedded in Chinese characters and integrate with other computational models. |
Copied to clipboard
| Challenge: | idioms and metaphors processing is a rapidly growing area in NLP, says dr. s. robertson . idiomatic idiomas are characteristic to all areas of human activity and to all types of discourse. |
| Approach: | This tutorial will provide attendees with a clear notion of idioms and metaphors . it will provide them with computational models of linguistic characteristics and methods . |
| Outcome: | This tutorial aims to provide attendees with a clear notion of the linguistic characteristics of idioms and metaphors . it outlines how to model idiomatic idiomes and their processing and what resources are available to support their use . |
Copied to clipboard
| Challenge: | a recent paper examines the relationship between semantics and pragmatics in language. |
| Approach: | They propose to develop computational models that leverage pragmatic knowledge in language . goal is to build better and more pragmatically-aware natural language generation and understanding systems . |
| Outcome: | The proposed models leverage pragmatic knowledge in language crucial to performing many NLP tasks correctly. |
Copied to clipboard
| Challenge: | a problem with natural language generation systems is the generation of tokens that are unrelated to the source input. |
| Approach: | They propose two new models to play the GuessWhat?! referential game . they propose to adapt the best visual processing models available to mitigate this issue . |
| Outcome: | The proposed models generate few hallucinations compared to other models available in the literature. |
Copied to clipboard
| Challenge: | Compounding is a prevalent word-formation process in Chinese morphology, where each character is bound and free when treated as a morpheme. |
| Approach: | They propose a model that learns non-linear relations between constituents and words and a character Jacobians model that describes character’s role in each word. |
| Outcome: | The proposed model predicts embeddings of real words from constituents but helps account for behavioral data of pseudowords. |
Copied to clipboard
| Challenge: | Humor is an essential but most fascinating element in personal communication. |
| Approach: | They propose a convolutional neural network with extensive filter size and filter number to increase the depth of networks. |
| Outcome: | The proposed model outperforms existing models on accuracy, precision and recall . the proposed model can learn to distinguish between humorous and nonhumorous texts . |
Copied to clipboard
| Challenge: | Prior research has focused on English language structures and multilingual contexts . however, there are several shortcomings with specialized sentence ordering models and advanced Large Language Models like GPT-4. |
| Approach: | They propose a multilingual sentence order task that extends SO to diverse narratives across 12 languages and code-switched texts. |
| Outcome: | The proposed task extends SO to diverse narratives across 12 languages, including challenging code-switched texts. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. |
| Approach: | They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets. |
| Outcome: | The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions. |
Copied to clipboard
| Challenge: | Polysynthetic languages are low-resource, lacking large scale annotated datasets needed to build and/or evaluate computational models. |
| Approach: | They propose to use linguistic priors to help with morphological segmentation and part-of-speech tagging tasks for Adyghe and Inuktitut . |
| Outcome: | The proposed methods improve morphological segmentation and part-of-speech tagging tasks on Adyghe and Inuktitut. |
Copied to clipboard
| Challenge: | Existing algorithms for dependency tree sampling have been proposed for sampling without replacement. |
| Approach: | They propose an algorithm that adapts the Wilson Reject algorithm for sampling without replacement and combines it with a Trie data structure. |
| Outcome: | The proposed method is efficient in the case of sampling without replacement from dependency graphs with random weights. |
Copied to clipboard
| Challenge: | Current approaches to model or simulate the acquisition of spoken language via grounding in perception are not generalizable to real-life situations that humans or adaptive artificial agents experience. |
| Approach: | They propose to use a dataset based on the children’s cartoon Peppa Pig to train a bi-modal architecture that learns aspects of the visual semantics of spoken language. |
| Outcome: | The proposed model learns to represent speech and visual data in a joint vector space. |
Copied to clipboard
| Challenge: | Existing work on automated essay scoring has focused on holistic scoring, which summarizes the quality of an essay with a single score. |
| Approach: | They present a corpus of essays simultaneously annotated with argument components, argument persuasiveness scores, and attributes of argument components that impact an argument’s persuasiveness. |
| Outcome: | The proposed corpus could trigger the development of novel computational models that provide useful feedback to students on why their arguments are (un)persuasive . |
Copied to clipboard
| Challenge: | a recent study has compared real and counterfactual word orders, but one functional pressure has been overlooked . a study of 10 typologically diverse languages shows that real word orders have greater uniformity than reverse word orders . |
| Approach: | They propose to test whether a pressure for UID may have influenced word order patterns cross-linguistically. |
| Outcome: | The proposed model shows that real orders have greater uniformity than reverse orders among SVO languages. |
Copied to clipboard
| Challenge: | Existing approaches to slang semantic variation do not account for the semantic variation of sling among different groups of users. |
| Approach: | They propose to use slang semantic variation models to trace the regional identity of a new emerging sling sense given its historical meanings. |
| Outcome: | The proposed models can predict regional identity of emerging slang word meanings from historical sling dictionary entries. |
Copied to clipboard
| Challenge: | Creating embodied, situated agents able to move in, communicate naturally about, and collaborate on human terms in the physical world has been a persisting goal in artificial intelligence (Winograd, 1972). |
| Approach: | They propose to use a 3D Minecraft dataset to model the beliefs of human partners in situ to enable theory of mind modeling in situated interactions. |
| Outcome: | The proposed model can be used to model human collaborative behaviors in the 3D virtual blocks world of Minecraft. |
Copied to clipboard
| Challenge: | linguistic constraints in dependency trees are not part of the definition of spanning trees. |
| Approach: | They propose to use a constraint that requires a single root to be incorporated into dependency tree sampling . they propose to reduce the asymptotic runtime of sampling k trees without replacement to O(kn3) |
| Outcome: | The proposed algorithms are asymptotically and practically more efficient . they reduce the runtime of the fastest algorithm for sampling with replacement to O(kn3) |
Copied to clipboard
| Challenge: | a novel approach to understanding narratives involves modelling the interaction between characters and actions . we propose role-playing games as a testbed for inferring interactions between characters in narratives . |
| Approach: | They propose role-playing games as a testbed for learning latent ties between characters and actions . they propose to combine character and action descriptions from online discussion forums . |
| Outcome: | The proposed model can capture interactions between characters and actions in narratives . it can predict actions better when character attributes are taken into account . |
Copied to clipboard
| Challenge: | Developing computational models that can produce contextually witty image descriptions is challenging because of the large corpus of sentences that are not available for large scale corpora. |
| Approach: | They propose to use linguistic wordplay, specifically puns, to generate witty image descriptions from large corpus of sentences or encode them via an encoder-decoder neural network architecture. |
| Outcome: | The proposed models perform better than baseline models using human data and show that they are slightly wittier than human-written witty descriptions. |
Copied to clipboard
| Challenge: | Discourse signals are often implicit, leaving it up to the interpreter to draw inferences . current discourse data and frameworks ignore the social aspect, expecting only a single ground truth . elisa f. and her team present a dataset with multiple and subjective interpretations of English conversation . |
| Approach: | They present a first discourse dataset with multiple and subjective interpretations of English conversation . they show disagreements are nuanced and require a deeper understanding of contextual factors . |
| Outcome: | The proposed dataset shows disagreements are nuanced and require deeper understanding of contextual factors. |
Copied to clipboard
| Challenge: | Prior studies on identifying the existence or the type of complaints focus on building automatic classification models for identifying complaints. |
| Approach: | They propose to measure the intensity of complaints from text using Best-Worst Scaling method to estimate the popularity of posts on social media. |
| Outcome: | The proposed model can estimate the popularity of complaints on social media with best-worst scaling (BWS) method. |
Copied to clipboard
| Challenge: | Recent work has focused on identifying narrative elements in personal stories texts, but this paper focuses on informational texts. |
| Approach: | They propose a novel NLP task for detecting narrative elements in raw text by adapting elements from the oral narrative theory of Labov and Waletzky and adding a new narrative element of their own. |
| Outcome: | The proposed scheme achieves an average F1 score of 0.77 and is better suited for informational texts than the oral narrative theory. |
Copied to clipboard
| Challenge: | A broad space separates its two constituent disciplines—natural language processing and social science—which has to date been sidestepped rather than filled by applying increasingly complex computational models to problems in social science research. |
| Approach: | They argue that computational text analysis lacks organizing principles and requires organizing methods to solve problems. |
| Outcome: | The proposed approach is based on a review of 60 papers on computational text analysis. |
Copied to clipboard
| Challenge: | a study aims to explore the role of speech pauses and gestures alone as predictors of audience reaction without other types of speech information. |
| Approach: | They analyze two speeches by Barack Obama and use them to predict audience reaction . they find that long pauses and co-speech gestures alone predict audience response . |
| Outcome: | The proposed models can predict audience reaction without other types of speech information. |
Copied to clipboard
| Challenge: | Mental health stigma prevents many individuals from receiving appropriate care, and social psychology studies have shown that mental health tends to be overlooked in men. |
| Approach: | They propose to use clinical psychology literature to curate prompts, then evaluate models’ propensity to generate gendered words. |
| Outcome: | The proposed framework captures stigma about gender in mental health and is more likely to predict female subjects than male in sentences about mental health conditions (32% vs. 19%), and this disparity is exacerbated for sentences that indicate treatment-seeking behavior. |
Copied to clipboard
| Challenge: | et al. argued that sentence co-occurrence probabilities should reflect entailment . but it is unclear whether probabilities predicted by neural LMs encode enanglement based on their theory . |
| Approach: | They propose a test that decodes entailment relations between natural sentences . they argue that the test that predicts a flipped test does not account for redundancy . |
| Outcome: | The proposed test can decode entailment relations between natural sentences, but not perfectly. |
Copied to clipboard
| Challenge: | Existing methods to detect online abuse focus on the more explicit forms of abuse . existing methods focus on detecting subtler forms of online abuse leaving them unnoticed . |
| Approach: | They propose a task to detect unpalatable questions using reddit data to implement a context-aware dataset and implement 'learning models' they hope future research will address subtle forms of abuse since harm passes unnoticed through existing detection systems. |
| Outcome: | The proposed task is based on a dataset of reddit users and a conversational context. |
Copied to clipboard
| Challenge: | This work describes IteraTeR: the first large-scale, multi-domain, edit-intention annotated corpus of iteratively revised text. |
| Approach: | They propose to annotate iteratively revised text using a multi-domain annotated corpus that generalizes to a variety of domains, edit intentions, revision depths, and granularities. |
| Outcome: | The proposed model improves automatic evaluations by integrating edit intentions with writing quality. |
Copied to clipboard
| Challenge: | Metaphor is a linguistic device in which a concept is expressed by mentioning another . Verbal MWEs are examples of non-literal language in which multiple words form a single unit of meaning. |
| Approach: | They propose to analyze the interplay between metaphor and multiword expressions processing by informing the model of the presence of MWEs. |
| Outcome: | The proposed architecture reach state-of-the-art on two established metaphor datasets. |
Copied to clipboard
| Challenge: | phonological hierarchies that predict coordinate constructions are often phonetically “natural” . a neural sequence labeling model can learn elaborate expressions in Hmong without using phonology information. |
| Approach: | They propose that coordinate compounds and elaborate expressions can be learned empirically by phonological hierarchies and a neural sequence labeling model can learn the ordering of elaborate expression in Hmong without using phonology. |
| Outcome: | The proposed models beat strong baselines for all three languages and learn hierarchies similar to those proposed by Mortensen. |
Copied to clipboard
| Challenge: | Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition. |
| Approach: | They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision. |
| Outcome: | The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision . |
Copied to clipboard
| Challenge: | Existing computational models of hate speech focus on a binary or multiclass classification task . a recent study shows an alarming 4.6% increase in hate speech in 2016 . |
| Approach: | They propose a task of deciphering hate symbols using the Urban Dictionary . they propose ciphers using Sequence-to-Sequence models and a Variational Decipher . |
| Outcome: | The proposed model can crack hate symbols based on context and generalize better to unseen symbols in a more challenging testing setting. |
Copied to clipboard
| Challenge: | Using computational models, the use of offensive language is pervasive in social media . a popular line of research is the study of machine learning classifiers to identify offensive content online . |
| Approach: | They analyze social media posts written by individuals with depression and those without . they train computational models to compare use of offensive language with depression detection . |
| Outcome: | The proposed models show that offensive language is more frequently used in the samples written by individuals with depression and those showing signs of depression. |
Copied to clipboard
| Challenge: | Existing methods to modify text understanding systems use only one sentence at a time . however, considering a larger context can improve performance for text understanding tasks. |
| Approach: | They propose to modify existing text data to insert out-of-context errors . they use a 2016 TEDTalk corpus to evaluate computational models for text understanding . |
| Outcome: | The proposed method targets real-world problems of transcription and translation systems by inserting authentic out-of-context errors. |
Copied to clipboard
| Challenge: | Mental health conditions remain underdiagnosed in many countries despite access to advanced medical care . a new approach to learn mood markers from mobile data is needed to improve accuracy and improve learning from typed text. |
| Approach: | They propose to use mobile data to learn mood markers without identifying users through personal or protected attributes. |
| Outcome: | The proposed model obfuscates user identities while remaining predictive . future directions include better models and pre-learning from typed text . |
Copied to clipboard
| Challenge: | linguists have suggested that some languages are "cooler" than others because of their contexts. |
| Approach: | They propose to omit plurality and definiteness markers in Chinese noun phrases . they build a corpus of Chinese NPs accompanied by its context . |
| Outcome: | The proposed model predicts the plurality and definiteness of Chinese noun phrases (NPs) it shows that speakers drop plurality markers very frequently, and that they are more likely to drop pronouns . |
Copied to clipboard
| Challenge: | Many studies on sentiment analysis focus on the fact that sentiment computations are compositional . linguistic utterances often do not adhere to strict patterns and can be surprising when looking at the individual words involved. |
| Approach: | They propose a method for obtaining non-compositionality ratings for phrases with respect to their sentiment . they also propose evaluating computational models for sentiment analysis using the rating resource . |
| Outcome: | The proposed method enables non-compositional ratings for phrases with respect to their sentiment . the results are compared with a new resource of ratings for 259 phrases . |
Copied to clipboard
| Challenge: | Visual illusions are a phenomenon that is often seen in human perception but are not always faithful to the physical world. |
| Approach: | They build a dataset containing five types of visual illusions and formulate four tasks to examine visual illusion in state-of-the-art VLMs. |
| Outcome: | The proposed dataset reveals that larger models are closer to human perception and more susceptible to visual illusions. |
Copied to clipboard
| Challenge: | Existing methods to overcome catastrophic forgetting in visual question answering models are inadequate, but have received little attention within natural language processing. |
| Approach: | They devise a set of linguistically-informed visual question answering tasks motivated by psycholinguistics and investigate impact of task difficulty on continual learning. |
| Outcome: | The proposed models differ in the types of questions they ask and show that task difficulty and order matter. |
Copied to clipboard
| Challenge: | Quantifiers are pervasive in NLU benchmarks and their occurrence at test time is associated with performance drops. |
| Approach: | They propose a generalized quantifier NLI task to quantify their contribution to the errors of NLU models. |
| Outcome: | The proposed model is based on a generalized quantifier theory and is compared with pre-trained models. |
Copied to clipboard
| Challenge: | Mars? - PragmatiCQA |
| Approach: | Mars? - The Paper . |
| Outcome: | The proposed dataset features 6873 QA pairs that explores pragmatic reasoning in conversations over a diverse set of topics. |
Copied to clipboard
| Challenge: | Existing work on automated essay scoring has focused on holistic scoring, but there is limited annotated corpus of essays with thesis strength scores. |
| Approach: | They propose a scoring rubric for persuasive essay quality and annotate corpus of essays with thesis strength scores. |
| Outcome: | The proposed scoring rubric could provide feedback to students on why essay gets thesis strength score . the rubric can be used to score persuasive essay quality, thesis strength, and organization . |
Copied to clipboard
| Challenge: | Existing work on argument quality (AQ) focuses on overall quality, but there is no large-scale theory-based corpus and corresponding computational models. |
| Approach: | They propose to use a large-scale English multi-domain argumentative writing corpus annotated with theory-based AQ scores to assess argument quality. |
| Outcome: | The proposed methods improve argument quality in three domains and can be used as strong baselines for future work. |
Copied to clipboard
| Challenge: | Existing studies on author profiling focus on age and gender, and use only English text. |
| Approach: | They propose to model author profiling from a Brazilian Portuguese corpus using standard gender and age prediction tasks and two less-known alternatives: predicting an author's degree of religiosity and IT background status. |
| Outcome: | The proposed tasks are based on a Brazilian Portuguese corpus and are compared with other languages and tasks. |
Copied to clipboard
| Challenge: | a lack of comprehensive datasets specifically annotated for hate instigating speech hinders research . lack of reliable models for hate triggering makes it difficult to apply off-the-shelf models to the problem. |
| Approach: | They propose to use a multilingual dataset to identify hate instigating speech . lack of comprehensive datasets specifically annotated for hate instigators hinders their work . |
| Outcome: | The proposed dataset identifies hate instigating speech across languages . lack of comprehensive datasets makes it difficult to train and evaluate models . |
Copied to clipboard
| Challenge: | Effective argumentation is essential towards a purposeful conversation with a satisfactory outcome. |
| Approach: | They propose a controllable neural argument generator capable of producing factual arguments from input facts and real-world concepts that can be explicitly controlled for stance and argument structure. |
| Outcome: | The proposed model produces factual arguments from input facts and real-world concepts that can be explicitly controlled for stance and argument structure using Walton’s argument scheme-based control codes. |
Copied to clipboard
| Challenge: | Referring Expression Generation (REG) lexical choice is the subtask that provides words to express an input meaning representation. |
| Approach: | They propose a personality-dependent lexical choice model for Referring Expression Generation (REG) that provides words to express a given input meaning representation. |
| Outcome: | The proposed model outperforms a standard lexicalisation model based on meaning-to-text mappings and personality information. |
Copied to clipboard
| Challenge: | Experimental results show that there is a significant performance gap between advanced models (72%) and humans (87%) Cloze datasets are convenient either to be generated automatically or by annotators. |
| Approach: | They propose to use a dataset to evaluate the performance of computational models through sentence prediction. |
| Outcome: | The proposed model fills up multiple blanks in a passage from a shared candidate set with distractors designed by English teachers. |
Copied to clipboard
| Challenge: | Existing data on compositionality of multi-word expressions is limited and only available for high resource languages. |
| Approach: | They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression . |
| Outcome: | The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality. |
Copied to clipboard
| Challenge: | lexical ambiguity is a problem for NLP, but few tasks evaluate its impact on human intuitions. |
| Approach: | They propose to use contextualized word embeddings to evaluate word meaning . they use a dataset of human relatedness judgments and human estimates of sense dominance . |
| Outcome: | The proposed model matches human intuitions with contextualized embeddings on 112 ambiguous words in context with 672 sentence pairs. |
Copied to clipboard
| Challenge: | Existing research on deception detection and fact checking conflates factual accuracy with truthfulness . a belief-based deception framework defines texts as deceptive when there is a mismatch between what people say and what they truly believe . |
| Approach: | They assess if presumed patterns of deception generalize to German language texts . they gauge the impact of deceptiveness on the downstream task of fact checking . |
| Outcome: | The proposed framework disentangles deception when there is a mismatch between what people say and what they truly believe . the proposed framework does not find any correlation with established cues of deception . |
Copied to clipboard
| Challenge: | Fallacies are used as seemingly valid arguments to support a position and persuade the audience about its validity. |
| Approach: | They propose to use instruction-based prompting to recognize 28 unique fallacies across datasets . they also analyze the effect of model size and prompt choice on model performance . |
| Outcome: | The proposed approach can recognize 28 unique fallacies across domains and genres. |
Copied to clipboard
| Challenge: | Readability assessment is the task of evaluating the reading difficulty of a given piece of text. |
| Approach: | They examine the common approaches used for automatic readability assessment and identify their shortcomings and some challenges for the future. |
| Outcome: | The proposed models are compared with existing models and are based on existing ones. |
Copied to clipboard
| Challenge: | Existing studies of Arabic dialects have focused on blogs and comments on online news sites, but data on other dialects are costly and limited. |
| Approach: | They present a dataset of > 1/4 billion tweets representing a wide range of Arabic dialects. |
| Outcome: | The dataset represents 29 major Arab cities from 10 Arab countries with varying dialects. |
Copied to clipboard
| Challenge: | Existing methods for image captioning do not guarantee consistent image-text relations . current models do not provide enough data for training robust captioning models . |
| Approach: | They use an annotation protocol specifically devised for capturing image–caption coherence relations to study image captioning. |
| Outcome: | The proposed protocol improves image captioning models with coherence relations . the dataset is large enough to alleviate content hallucinations, the authors show . |
Copied to clipboard
| Challenge: | Existing algorithms and tools for sentiment analysis are lacking in dealing with Arabic metaphorical expressions. |
| Approach: | They propose to use Arabic metaphors in automatic Arabic sentiment analysis to examine the performance of a state-of-art Arabic sentiment tool on metaphors. |
| Outcome: | The proposed model outperforms the state-of-the-art sentiment analysis tool on metaphors and gain a deeper insight into the issue. |
Copied to clipboard
| Challenge: | Politeness principles play a central role in shaping human interaction. |
| Approach: | They propose a generalized framework for modeling face acts in persuasion conversations using an annotated corpus and computational models. |
| Outcome: | The proposed framework reveals differences in face act utilization between asymmetric roles in persuasion conversations and predicts key conversational outcome. |
Copied to clipboard
| Challenge: | Existing annotations for irony are difficult, and the recognition of it is difficult due to its polarity. |
| Approach: | They propose a fine-grained annotation scheme centered on irony that highlights the tokens responsible for its activation and their morpho-syntactic features. |
| Outcome: | The proposed scheme highlights the tokens responsible for irony activation and their morpho-syntactic features. |
Copied to clipboard
| Challenge: | a gap in the literature on offensive language has been addressed with studies on Spanish, Hindi, and German. |
| Approach: | They present a Greek annotated dataset for offensive language identification . it contains 4,779 tweets annotating offensive and not offensive posts from Twitter . they evaluate several computational models trained and tested on the dataset . |
| Outcome: | The proposed dataset contains 4,779 tweets annotated as offensive and not offensive . the authors show that the proposed dataset is similar to the OLID dataset for English . |
Copied to clipboard
| Challenge: | Winograd schemas are well-established tools for evaluating coreference resolution and commonsense reasoning capabilities of computational models. |
| Approach: | They present a dataset of German, French, and Russian schemas aligned with their English counterparts. |
| Outcome: | The proposed model improves in English and German, while the model improve in other languages. |
Copied to clipboard
| Challenge: | Using computational models as pedagogical tools is becoming increasingly popular, but how effective can these models adapt as teachers to students of different types? |
| Approach: | They propose a suite of models and evaluation methods that combine Bayesian student models and AToM to evaluate adaptive teaching methods. |
| Outcome: | The proposed models outperform LLM-based and standard Bayesian teaching methods in the evaluation of simulated students across three learning domains. |
Copied to clipboard
| Challenge: | comparative method allows linguists to infer protoforms from their reflexes based on sound change . authors argue that this approach ignores one of the most important aspects of the comparative approach . |
| Approach: | They propose a comparative method that allows linguists to infer protoforms from their reflexes . they propose to use a system where candidate protoform from a reconstruction model are reranked by a reflex prediction model. |
| Outcome: | The comparative method surpasses state-of-the-art methods on Chinese and Romance datasets. |
Copied to clipboard
| Challenge: | a new multimodal dataset of stand-up comedies is proposed to improve humor detection . the dataset is the biggest available for this type of task, and the most diverse . |
| Approach: | They propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors. |
| Outcome: | The proposed method improves existing models of humor detection by using audio speech recognition errors. |
Copied to clipboard
| Challenge: | Existing studies on MT evaluation characterize quality of output with a single number . a recent advancement in MT technologies has enabled higher-quality, more nuanced translations . |
| Approach: | They propose a 1200-sentence MQM evaluation benchmark for English-Korean and a reference-free QE setup to evaluate the quality of the translations. |
| Outcome: | The proposed model outperforms the existing model in style and accuracy. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) are gaining popularity as partial solutions to the “symbol grounding problem” faced by language models trained on text alone. |
| Approach: | They propose to use multimodal large language models to integrate linguistic representations with data from other modalities to investigate whether they are integrated into a model. |
| Outcome: | The proposed models are sensitive to visual features like object shape when it is implied by a verbal description of an event. |
Copied to clipboard
| Challenge: | a number of languages are used in online conversations, resulting in code-mixing . the problem is largely unexplored due to the lack of annotated data and noise . |
| Approach: | They propose a robust perturbation-based joint-training model that learns to handle noise in code-mixed text by parameter sharing across clean and noisy words. |
| Outcome: | The proposed model learns to handle noise in the real-world code-mixed text by parameter sharing across clean and noisy words. |
Copied to clipboard
| Challenge: | Using minimal pairs and surprisal-based measures, we assess whether large language models exhibit systematic biases toward non-local antecedents in logophoric contexts. |
| Approach: | They examine large language models’ sensitivity to four logophoric cues known to license long-distance binding of the reflexive ziji . |
| Outcome: | The proposed model families show that they exhibit above-chance sensitivity to all four cues, while lexically anchored cue are more robustly captured than discourse-level cue. |